Papers with computational analysis
OpenFraming: Open-sourced Tool for Computational Framing Analysis of Multilingual Data (2021.emnlp-demo)
Copied to clipboard
Vibhu Bhatia, Vidya Prasad Akavoor, Sejin Paik, Lei Guo, Mona Jalal, Alyssa Smith, David Assefa Tofu, Edward Edberg Halim, Yimeng Sun, Margrit Betke, Prakash Ishwar, Derry Tanti Wijaya
| Challenge: | Existing frameworks for analyzing frames in multilingual text documents are available online and via an API. |
| Approach: | They propose a web-based system for analyzing frames in multilingual text documents . framework combines unsupervised and supervised machine learning and leverages a state-of-the-art multilingual language model . |
| Outcome: | The proposed framework can significantly improve frame prediction performance while requiring a small sample of manual annotations. |
SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are being used in urban planning but there is concern that they reproduce or amplify such biases. |
| Approach: | They propose a framework to evaluate spatial gender bias in large language models . they use a taxonomy of 62 urban micro-spaces, a prompt library and three diagnostic layers . |
| Outcome: | The proposed framework identifies structured gender-space associations that go beyond the public-private divide, forming nuanced micro-level mappings. |
On the Impact of Temporal Representations on Metaphor Detection (2022.lrec-1)
Copied to clipboard
| Challenge: | State-of-the-art approaches for metaphor detection compare their literal - or core - meaning and their contextual meaning using neural networks. |
| Approach: | They propose to use temporal and static word embeddings to account for different representations of literal meanings to examine metaphor detection tasks. |
| Outcome: | The proposed method outperforms static methods but may provide representations of the core meaning of the metaphor too close to their contextual meaning, causing confusion. |
The Timing of Lexical Memory Retrievals in Language Production (N18-1)
Copied to clipboard
| Challenge: | In a large-scale observational study of a spoken corpus, we find that language production at a time point preceding a word is sped up or slowed down depending on activation of that word. |
| Approach: | They propose a cognitive model of fluency in which lexical memory retrievals may explain some of the variability in speech rates. |
| Outcome: | The proposed model predicts that language production is sped up or slowed down depending on activation of a word . |
A Workflow for HTR-Postprocessing, Labeling and Classifying Diachronic and Regional Variation in Pre-Modern Slavic Texts (2024.lrec-main)
Copied to clipboard
Piroska Lendvai, Maarten van Gompel, Anna Jouravel, Elena Renje, Uwe Reichel, Achim Rabus, Eckhart Arnold
| Challenge: | a workflow for classifying diachronic and regional language variation in medieval texts is currently being developed . the workflow is generic or language-agnostic, but can be applied to other historical languages as well. |
| Approach: | They propose a workflow for classifying diachronic and regional language variation in medieval texts . they use handwritten text recognition and manual transcription to obtain the data . |
| Outcome: | The proposed workflow covers HTR-postprocessing, annotating and classifying medieval texts . it is accessible to humanists with limited experience in research data infrastructures, analysis or NLP . |
A Taxonomy of Empathetic Questions in Social Dialogs (2022.acl-long)
Copied to clipboard
| Challenge: | Current dialog generation approaches do not model effective question-asking due to the lack of a taxonomy of questions and their purpose in social chitchat. |
| Approach: | They propose to model questions' ability to capture communicative acts and their emotion-regulation intents by annotating a large dataset with established labels. |
| Outcome: | The proposed model can be used to generate labels for the EmpatheticDialogues dataset and to further improve the existing models. |
Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to analyze moral reasoning are discordant and lack cohesion, focusing on isolated aspects of the process. |
| Approach: | They propose a unified dataset that integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, and captures diverse socio-cultural contexts. |
| Outcome: | The proposed dataset integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, along with annotators’ moral and cultural profiles. |
HisDoc-OCR: Restoring Visual Grounding in MLLMs for Chinese Historical Document OCR (2026.findings-acl)
Copied to clipboard
| Challenge: | Despite multimodal large language models' strong performance on modern document OCR, their application to historical Chinese texts suffers from severe hallucinations, character fabrication, uncontrolled repetition, and semantic drift. |
| Approach: | They propose a multimodal large language model which restores visual grounding through three synergistic strategies: Layout Injection, First-Occurrence Boost, Self-Distilled Attention Focusing and HisDoc-OCR. |
| Outcome: | The proposed model outperforms general-purpose and OCR-specific models on Chinese historical documents. |
A Corpus of Natural Multimodal Spatial Scene Descriptions (L18-1)
Copied to clipboard
| Challenge: | Existing work on multimodal spatial descriptions combines speech and hand gestures to form a corpus of multimodal descriptions. |
| Approach: | They present a corpus of multimodal spatial descriptions as commonly occurring in route giving tasks. |
| Outcome: | The proposed corpus of multimodal spatial descriptions is more amenable to computational analysis and useable for learning natural computer interfaces. |
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia (2024.acl-long)
Copied to clipboard
Giovanni Monea, Maxime Peyrard, Martin Josifoski, Vishrav Chaudhary, Jason Eisner, Emre Kiciman, Hamid Palangi, Barun Patra, Robert West
| Challenge: | Large language models (LLMs) have an impressive ability to draw on novel information supplied in their context, yet the mechanisms underlying contextual grounding remain unknown. |
| Approach: | They propose a method to study grounding abilities using a counterfactual dataset constructed to clash with a model's parametric knowledge using Fakepedia. |
| Outcome: | The proposed method evaluates grounding abilities when the internal parametric knowledge clashes with the contextual information. |
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)
Copied to clipboard
| Challenge: | Dialectal Arabic datasets embody a range of domain, dialect, and quality. |
| Approach: | They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects. |
| Outcome: | The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing. |
Investigating Sports Commentator Bias within a Large Corpus of American Football Broadcasts (D19-1)
Copied to clipboard
| Challenge: | a recent study shows that sports broadcasters build drama into play-by-play commentary by building team and player narratives through subjective analyses and anecdotes. |
| Approach: | They use FOOTBALL to examine racial bias in sports commentary . they identify major confounding factors for researchers examining rraecial bias . |
| Outcome: | The proposed dataset supports previous social science studies on commentator bias . it contains 1,455 broadcast football transcripts annotated with 250K player mentions and racial metadata . |
A Dual-View Analysis of Multiple Languages in Colonial Newspapers (2026.findings-acl)
Copied to clipboard
Zhan Su, Xiaoya Chen, Fengran Mo, Ida L. Vos, Prayag Tiwari, Yazhou Zhang, Qian Zheng, Natália da Silva Perez
| Challenge: | Historical newspapers from the colonial period offer valuable evidence of how racializing language evolved over time. |
| Approach: | They propose a contextual question answering and visual question answering task from colonial newspapers . they propose linguistic training for temporal word embedding with a compass to study racialization . |
| Outcome: | The proposed tasks are limited for low-resource tasks, the authors show . the authors compare the results of two QA pairs from colonial newspapers to a compass . |
N-CORE: N-View Consistency Regularization for Disentangled Representation Learning in Nonverbal Vocalizations (2025.emnlp-main)
Copied to clipboard
| Challenge: | Nonverbal vocalizations are an essential component of human communication, conveying rich information without linguistic content. |
| Approach: | They propose a backbone-agnostic framework to disentangle emotion and speaker information from nonverbal vocalizations by leveraging N views of audio samples to learn invariance to specific transformations. |
| Outcome: | The proposed framework achieves competitive performance compared to state-of-the-art methods on the VIVAE, ReCANVo, and ReCANVO-Balanced datasets. |